Papers with data generation framework
What is it? Towards a Generalizable Native American Language Identification System (2025.naacl-srw)
Copied to clipboard
| Challenge: | Despite their cultural and historical significance, Native American languages remain unsupported by major commercial language identification systems. |
| Approach: | They propose to curate linguistic resources across all Native American languages for robust training and tailor data augmentation to generate synthetic yet linguistically coherent training samples. |
| Outcome: | The proposed system would be generalizable across all Native American languages . it would also generate coherent training samples for low-resource languages based on Plains Apache . |
UnSeenTimeQA: Time-Sensitive Question-Answering Beyond LLMs’ Memorization (2025.acl-long)
Copied to clipboard
Md Nayem Uddin, Amir Saeidi, Divij Handa, Agastya Seth, Tran Cao Son, Eduardo Blanco, Steven Corman, Chitta Baral
| Challenge: | UnSeenTimeQA is a data contamination-free time-sensitive question-answering benchmark. |
| Approach: | They propose a data contamination-free time-sensitive question-answering benchmark that avoids web-searchable queries grounded in the real world. |
| Outcome: | The proposed benchmark avoids web-searchable queries grounded in the real world and enables on-demand generation of new samples, mitigating the risk of data leakage. |
SynthDST: Synthetic Data is All You Need for Few-Shot Dialog State Tracking (2024.eacl-long)
Copied to clipboard
| Challenge: | In-context learning with Large Language Models (LLMs) is a promising avenue of research in Dialog State Tracking (DST). |
| Approach: | They propose a data generation framework tailored for Dialog State Tracking that uses large language models to synthesize natural, coherent, and free-flowing dialogues with DST annotations. |
| Outcome: | The proposed framework improves joint goal accuracy by 4-5% over the zero-shot baseline on MultiWOZ 2.1 and 2.4. |
Mitigating Gender Bias via Fostering Exploratory Thinking in LLMs (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large Language Models often exhibit gender bias, resulting in unequal treatment of male and female subjects across contexts. |
| Approach: | They propose a framework that encourages exploratory thinking in large language models . the framework generates story pairs featuring male and female protagonists in structurally identical scenarios . |
| Outcome: | The proposed framework reduces gender bias while preserving or even enhancing general model capabilities. |
Towards Better Hierarchical Text Classification with Data Generation (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to improve hierarchical text classification are expensive and lack high-quality labeled data. |
| Approach: | They propose a hierarchical text classification framework that can achieve both label controllability and text diversity by extracting high-quality hierarchic label information. |
| Outcome: | The proposed method can achieve label controllability and text diversity by extracting high-quality hierarchical label information. |
MP2D: An Automated Topic Shift Dialogue Generation Framework Leveraging Knowledge Graphs (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to manage topic shifts within on-topic dialogues are limited in their ability to generate training datasets. |
| Approach: | They propose a data generation framework that automatically generates conversational question-answering datasets with natural topic transitions by leveraging relationships between entities in a knowledge graph. |
| Outcome: | The proposed framework generates conversational question-answering datasets with natural topic transitions and proves its effectiveness in generating dialogues with topic shifts. |
MDCure: A Scalable Pipeline for Multi-Document Instruction-Following (2025.acl-long)
Copied to clipboard
| Challenge: | Multi-document (MD) processing is crucial for LLMs to handle real-world tasks such as summarization and question-answering across large sets of documents. |
| Approach: | They propose a framework that generates high-quality synthetic MD instruction data over sets of articles via targeted prompts. |
| Outcome: | MDCure generates high-quality synthetic MD instruction data over sets of articles . evaluations show it improves over pre-trained models by up to 75.1% . |
PhaseMI: A Motivational Interviewing Dataset for Enhancing Phase Progression in LLM-based Counseling (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing MI datasets do not explicitly model structured progression of MI phases, which is essential for effective and goal-oriented counseling. |
| Approach: | They propose a phase-structured MI dataset with a data generation framework that employs therapist, client, and supervisor LLMs to explicitly control phase transitions. |
| Outcome: | The proposed model achieves 12.3% better coverage of MI phases, 37.6% in guiding, and 61.1% in choosing. |